Back

BMJ Health & Care Informatics

BMJ

All preprints, ranked by how well they match BMJ Health & Care Informatics's content profile, based on 15 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Can ChatGPT give holistic and accurate patient-centred information to oncology patients? A mixed-methods evaluation with stakeholders

Sun, M.; Reiter, E.; Murchie, P.; Kiltie, A. E.; Ramsay, G.; Duncan, L.; Adam, R.

2026-02-03 primary care research 10.64898/2026.02.02.26345346 medRxiv
Top 0.1%
38.8%
Show abstract

ObjectiveMore people than ever before are living with cancer. Patient education is a core component of cancer care, and patients are increasingly using large language models (LLMs), such as ChatGPT, for advice. The objectives of this study were to evaluate the ability of ChatGPT to explain specialist cancer care records (multidisciplinary team (MDT) meeting reports) to patients and to understand key stakeholders views and opinions about the technology. MethodsSix simulated MDT meeting reports were created by cancer clinicians. MDT reports and 184 realistic patient-centred queries were input into ChatGPT4.0 web version. We conducted a mixed-methods study combining qualitative analysis with exploratory quantitative components to evaluate ChatGPTs responses. The study consisted of three stages: (1) Clinician sense-checking, (2) Clinical and non-clinical annotation, (3) focus groups (including cancer patients, caregivers, computer scientists, and clinicians). ResultsChatGPT was able to summarise complex oncology information into simpler language, to provide definitions of complex terms and to answer questions about clinical care. However, clinician sense-checking identified problems with accuracy, language and content. In clinician annotation, 92.6% of ChatGPTs responses were judged problematic. Across all evaluation methods, six recurring themes were identified: accuracy, language, trust, content, personalisation and integration challenges. Patients and clinicians found the summaries and definitions useful; however, the responses were not tailored to the individual patient or to what the report might mean for them. ConclusionThis study highlights current challenges in using LLMs to explain complex cancer diagnoses and treatment records, including inaccurate information, inappropriate language, limited personalisation, AI distrust and challenges in integrating LLMs into clinical workflow. Understanding of the limitations is crucial for clinicians, patients, computer scientists and policy makers. The issues should be addressed before deploying LLMs in clinical settings.

2
Priorities for AI Education: Clinicians' Perspectives

Jeffrey, M.; Auyoung, E.; Pak, D.

2025-07-30 health informatics 10.1101/2025.07.29.25330662 medRxiv
Top 0.1%
38.4%
Show abstract

ObjectiveEducating clinicians about Artificial Intelligence (AI) is an urgent need(1) as the UK General Medical Council (GMC) places liability with practitioners(2) and the EU AI Act with employers to provide appropriate training(3), but also because AI, like any tool, requires training to use safely. NHSE Capability Framework provides guidance(4), but frontline clinicians perspectives are unknown so we sought to identify their priorities. Methods and AnalysisUsing iterative interviews with residents, educators and experts we synthesised 10 contextualised AI-related problem statements. We surveyed residents and consultant-educators in the East of England, who rated their confidence and importance. Participants also ranked their preferred learning modality. ResultsWe received 299 responses. Clinicians priorities, defined by high importance (I) and low confidence (C), were: understanding liability implications (I: 40%; C: 1.82/5), determining appropriate levels of confidence in AI algorithms (I: 36.5%; C: 1.98/5), and mitigating security and privacy risks (I: 34%; C: 1.68). Confidence was low (mean 20, range 10-50), with no significant difference between educators and residents. Residents preferred integration of training into regional teaching, while consultant-educators favoured webinars. ConclusionOur findings show that clinicians prioritise practical concerns, such as liability and determining confidence in algorithmic outputs. In contrast, critical appraisal and explaining AI to patients were deprioritised, despite their relevance to clinical safety. This study enhances the NHSE Capability Framework by contextualizing AI-related capabilities for clinicians as users and identifying priorities with which to develop scalable training. Key MessagesO_ST_ABSWhat is already know on this topicC_ST_ABSWhile clinicians face legal accountability for their use of AI in healthcare(2,3,5), there remains no standardised educational pathway to support them in acquiring the necessary skills. Although expert-informed capability frameworks exist(6), they are necessarily broad and lack operational clarity for day-to-day clinical roles. What this study addsThis study translates 31 AI-related capabilities from the NHSE DART-Ed Capability Framework(6) into 10 concise AI learning needs for clinicians of the user archetype through iterative interviews with residents, educators and AI experts. A regional survey with 299 responses from residents and educators highlights practical concerns such as liability and determining appropriate confidence in AI algorithms as learners priorities, whilst critical appraisal and explaining AI to patients were deprioritised despite their relevance to clinical safety. How this study might affect research, practice or policyThe educational priorities of clinicians as users of AI identified in this study provides engaging, curriculum-ready content mapped to the user archetype of the DART-Ed framework, which can be adapted to role and task-specific educational activities.

3
Explainable AI and public reactions to AI-involved adverse diagnostic events: a vignette study

Choi, J.; Kim, Y. J.; Lyu, P.; Luan, Y. L.; Toh, S. M.

2026-06-02 health informatics 10.64898/2026.05.26.26353870 medRxiv
Top 0.1%
26.6%
Show abstract

Artificial intelligence (AI) is increasingly incorporated into diagnostic decision-making, raising questions about physician responsibility following AI-involved adverse diagnostic events. Explainable AI (XAI) has been proposed to improve transparency and trust, but its influence on public reactions remains unclear. In a randomised vignette-based experiment, 652 adults from the United States and United Kingdom were assigned to one of six conditions in a 3 (diagnostic source: AI alone, human radiologist alone, or human-AI collaboration) x 2 (explanation: present or absent) between-subjects design. Participants read a scenario in which a chest X-ray was initially interpreted as normal but lung cancer was diagnosed five months later, indicating that the original interpretation had missed the cancer. In explanation conditions, participants received additional information about how the diagnosis had been reached, including AI heatmap-based explanations in the AI conditions. Participants rated radiologist responsibility, likelihood of complaint, and intention to pursue legal action. Among 652 participants (mean age 42.2 years; 50.2% female), responsibility ratings were significantly lower when AI alone made the diagnostic decision (mean 4.73, 95% CI 4.53-4.93) compared with human-only decision-making (5.78, 95% CI 5.59-5.98; p<0.001) and human-AI collaboration (5.54, 95% CI 5.34-5.74; p<0.001). Complaint likelihood showed a similar pattern. Intentions to pursue legal action followed the same directional trend but were marginally significant. Neither explanations nor explanation-by-source interactions were associated with outcome measures. These findings suggest that the public expects physicians to remain accountable when AI is involved in diagnostic decision-making, particularly in collaborative settings. Providing explanatory information about how AI systems reach decisions may be insufficient to change perceptions of physician responsibility following adverse diagnostic events.

4
Exploring the Interpretability of AI Decision Support Systems for Surgical Anatomy Recognition

Khan, D. Z.; Adams, T.; Wijekoon, A.; Ramirez Herrera, R.; Bano, S.; McCulloch, P.; Stoyanov, D.; Clarkson, M. J.; Costanza, E.; Blandford, A.; Marcus, H.; CARES Evaluation Group,

2026-06-03 surgery 10.64898/2026.06.02.26354729 medRxiv
Top 0.1%
22.8%
Show abstract

Artificial intelligence (AI) decision support systems for surgery hold promise but face barriers to adoption, particularly around the interpretability of their outputs. We conducted an international cross-sectional survey of 47 neurosurgeons to evaluate perspectives on literature-derived explanation techniques for AI-generated anatomical segmentations, using endoscopic pituitary surgery as a high-risk exemplar. Participants ranked certainty scores, certainty maps, saliency maps, scene similarity scores, and nearest-neighbour illustrations, and rated them using a modified Explanation Satisfaction Scale alongside free-text feedback. Certainty-based techniques were consistently ranked and rated highest for interpretability - valued for aligning with surgical decision-making by conveying confidence (via scores) and anatomical boundaries (via maps). Saliency- and similarity-based methods were judged less clinically relevant and better suited to educational settings. Certainty-based explanations, therefore, appear most acceptable to surgeons for clinical integration of decision support systems, though their impact on AI acceptability, trust calibration, and performance requires prospective evaluation across surgical domains.

5
PREFER-IT: A transdisciplinary co-created framework to realise inclusive medical AI

Pita Ferreira, P.; Soriano Longaron, S.; Bouisaghouane, W.; Goris, J.; H. Hoekman, A.; Markos, B.; Maus, B.; Pozzi, G.; Hasan, H.; Kalinauskaite, I.; Stunt, J.; D. Kist, J.; van der Elst, J.; Maguet, K.; Ziegfeld, L.; Cuypers, M.; Milota, M.; Habets, M.; Colombo, S.; Petric, S.; Groefsema, S.; Warmelink, S.; Daae, E.; Briganti, G.; Vajda, I.; Valdenegro-Toro, M.; Braun, M.; Jeekel, P.; Goosen, S.; Schepel, A.; Ester, L.; Kuzee, R.; de Klerk, S.; Lamoth, C.; Ballard, L.; Plantinga, M.

2025-11-06 health informatics 10.1101/2025.11.03.25339443 medRxiv
Top 0.1%
21.8%
Show abstract

Artificial intelligence (AI) in healthcare holds transformative potential but risks exacerbating existing health disparities if inclusivity is not explicitly accounted for. This study addresses the disconnected discussions on inclusive medical AI by developing a comprehensive framework, PREFER-IT. This framework is based on the outcomes of a five-day transdisciplinary co-creation workshop that involved 37 experts from diverse backgrounds, including healthcare, ethics, law, social sciences, AI, and patient advocacy. For this workshop, we used design thinking and participatory methodologies to develop a framework for realising inclusive medical AI. We identified three key challenges for realising inclusive medical AI: integrating the lived experiences and stakeholder voices across the AI lifecycle, designing data collection practices that promote fairness and prevent inequalities, and fostering regulatory frameworks to uphold human rights and promote inclusivity. The analysis of participants perspectives informed the development of eight key thematic clusters of PREFER-IT: Participatory and co-design approaches (P), Representative and diverse data (R), Education and digital literacy (E), Fairness (F), Ethical and legal accountability (E), Real-world validation and feedback (R), Inclusive communication (I), and Technical interoperability (T). These elements were mapped across structural layers of AI (humans, data, system, process, and governance) and the AI lifecycle to guide inclusive design, development, validation, implementation, monitoring, and governance. This framework fosters stakeholder engagement and systemic change, positioning inclusion as a guiding principle in practice. PREFER-IT offers a practical and conceptual contribution for how to include ethical, legal and societal aspects when aiming to foster responsible and inclusive AI in healthcare. Author SummaryArtificial intelligence (AI) is being used more and more in healthcare to improve diagnosis, treatment, and personalised care. However, if not designed carefully, these technologies can unintentionally increase existing inequalities and exclude certain groups from their benefits. In our study, we brought together experts from healthcare, ethics, law, social sciences, and patient advocacy to explore how AI in medicine can be made more inclusive. Over five days, we worked together to identify key issues and come up with practical solutions. We focused on three main areas: 1) Ensuring diverse voices are heard during the development of AI tools; 2) Making data collection fair and representative; and 3) Creating regulations that protect human rights. From the discussions of the workshop, we created the PREFER-IT framework, which outlines eight key principles for inclusive AI: O_LIParticipatory and co-design approaches C_LIO_LIRepresentative and diverse data C_LIO_LIEducation and digital literacy C_LIO_LIFairness C_LIO_LIEthical and legal accountability C_LIO_LIReal-world validation and feedback C_LIO_LIInclusive communication C_LIO_LITechnical interoperability C_LI This framework helps guide developers, policymakers, and healthcare professionals in creating AI systems that are not only effective but also fair and respectful of all users. Our work emphasises the importance of involving patients and communities in shaping the future of AI.

6
Patient-centric radiology: Utilising large language models (LLMs) to improve patient communication and education

Yip, A.; Craig, G.; White, N. M.; Cortes-Ramirez, J.; Shaw, K.; Reddy, S.

2026-02-25 health informatics 10.64898/2026.02.23.26346923 medRxiv
Top 0.1%
19.2%
Show abstract

PurposeTo evaluate whether large language models (LLMs) can enhance clinician-patient communication by simplifying radiology reports to improve patient readability and comprehension. MethodsA randomised controlled trial was conducted at a single healthcare service for patients undergoing X-ray, ultrasound or computed tomography between May 2025 and June 2025. Participants were randomised in a 1:1 ratio to receive either (1) the formal radiology report only or (2) the formal radiology report and an LLM-simplified version. Readability scores, including the Simple Measure of Gobbledygook, Automated Readability Index, Flesch Reading Ease, and Flesch-Kincaid grade level, were calculated for both reports. Statistical analysis of patient readability and comprehension levels, factual accuracy and hallucination rates for LLMs was assessed using a combination of binary and 5-point Likert scales, open-ended survey questions, and independent review by two radiologists. Results59/120 patients were randomised to receive both the formal and LLM-simplified radiology reports. Readability of LLM-simplified reports significantly improved with the reading level required for formal reports equivalent to a university-standard (11th-13th grade) compared to a middle-school standard (5th-9th grade) for simplified reports (rank biserial correlation=0.83, p<0.001). Patients with both reports demonstrated a significantly greater comprehension level, with 95% reporting an understanding level greater than 50%, compared with 46% without the simplified report (rank biserial correlation = 0.67, p < 0.001). All LLM-simplified reports were considered at least somewhat accurate with a minimal hallucination rate of 1.7%. Importantly, no hallucinations resulted in potential patient harm. 118/120 (98.3%) patients expressed interest in simplified radiology reports to be included in future clinical practice. ConclusionThis study provides evidence that LLMs can simplify radiology reports to an accessible level of readability with minimal hallucination. LLMs improve both ease of readability and comprehension of radiology reports for patients. Therefore, the rapid advancement of LLMs shows strong potential in enhancing patient-radiologist communication as patient access to electronic health records is increasingly adopted. HighlightsO_LIRadiology reports can be complex and difficult for patients to read and interpret C_LIO_LIStrong patient demand exists for simplified radiology reports C_LIO_LILarge language models (LLMs) such as GPT-4o show promise in simplifying radiology reports C_LIO_LILLMs credibly simplify radiology reports with minimal hallucination rates C_LIO_LILLMs improve both patient readability and comprehension of radiology reports C_LI

7
Calibrating trust in AI-assisted pituitary surgery

Hudson, G. R.; Khan, D. Z.; Fayez, F.; Bhatia, S.; Bano, S.; Costanza, E.; Blandford, A.; Stoyanov, D.; McCulloch, P.; Marcus, H. J.; University College London Collaborators,

2026-06-04 surgery 10.64898/2026.06.02.26354735 medRxiv
Top 0.1%
19.1%
Show abstract

Background: Endoscopic endonasal transsphenoidal surgery (EETS) requires navigation around neurocritical anatomy. Today, artificial intelligence clinical decision support systems (AI-CDSSs) can orientate surgeons, but clinician trust in AI remains unclear, limiting safe deployment. This study evaluates how modifiable design affects trust and performance in a real-world pituitary surgery AI-CDSS. Method: Online, 70 clinicians with pituitary surgery experience were randomised evenly to a Basic or Enhanced AI-CDSS which outline the sella on EETS operative video. The Enhanced group additionally received explanation of the model and previous publications, alongside confidence labels depicting outline reliability. Both groups annotated the sella on six video clips, first alone then with the optional AI-CDSS. Clips were ordered by declining AI performance, except for the final clip. Self-reported trust was measured using a 1-7 scale after each annotation, and performance was the DICE overlap between user annotations and the ground truth. Comparisons used Mann-Whitney U and permutation analysis. Results: Sixty-four participants (91%) finished the exercise (31 Basic, 33 Enhanced). When AI performed best, median trust was 5.00 in both arms (U=559, p=.521). However, when AI performed worst, trust was significantly lower for the Enhanced group (3.00 vs 3.67, U=668, p=.035), sustained in the final clip (3.67 vs 4.33 U=687, p=.019). User performance improved with the AI-CDSS, but with no significant difference between the groups on the best or worst AI performing clips. Nevertheless, for the best AI, senior clinicians had higher median performance in the Enhanced group (0.95 vs 0.90, U=75, p=.066). There was also less dispersion in the Enhanced group when AI was inaccurate (IQR: 0.07 vs 0.21, p=.004). Conclusion: Interface design can improve trust calibration in a surgical AI-CDSS and may increment performance in seniors when AI is accurate, and consistency when AI is inaccurate. In future, these features may form important safety checks during translation to the operating room.

8
Patient Attitudes Toward Artificial Intelligence in Cancer Care: A Scoping Review

Hilbers, D.; Nekain, N.; Bates, A.; Nunez, J.-J.

2025-03-17 health informatics 10.1101/2025.03.15.25324029 medRxiv
Top 0.1%
18.8%
Show abstract

PURPOSETo synthesize existing literature on patient attitudes toward AI in cancer care and identify knowledge gaps that can inform future research and clinical implementation. DESIGNA scoping review was conducted following PRISMA-ScR guidelines. MEDLINE, EMBASE, PsycINFO, and CINAHL were searched for peer-reviewed primary research studies published until February 1, 2025. The Population-Concept-Context framework guided study selection, focusing on adult patients with cancer and their attitudes toward AI. Studies with quantitative or qualitative data were included. Two independent reviewers screened studies, with a third resolving disagreements. Data were synthesized into tabular and narrative summaries. RESULTSOur search yielded 1,240 citations, of which 19 studies met the inclusion criteria, representing 2,114 patients with cancer across 15 countries. Most studies used quantitative methods (n=9) such as questionnaires or surveys. The most studied cancers were prostate, melanoma, breast, and colorectal cancer. While patients with cancer generally supported AI when used as a physician-guided tool, concerns about depersonalization, treatment bias, and data security highlighted challenges in implementation. Trust in AI was shaped by physician endorsement and patient familiarity, with greater trust when AI was physician-guided. Geographic differences were observed, with greater AI acceptance in Asia, while skepticism was more prevalent in North America and Europe. Additionally, patients with metastatic cancer were underrepresented, limiting insights into AI perceptions in this population. CONCLUSIONThis scoping review provides the first synthesis of patient attitudes toward AI across all cancer types and highlights concerns unique to patients with cancer. Clinicians can use these findings to enhance patient acceptance of AI by positioning it as a physician-guided tool and ensuring its integration aligns with patient values and expectations.

9
Evaluating reasoning LLMs' potential to perpetuate racial and gender disease stereotypes in healthcare

Docking, J. J.; Li, L. X.; Menz, B. D.; Bacchi, S.; Hopkins, A. M.; Sorich, M. J.

2025-08-07 health informatics 10.1101/2025.08.05.25333007 medRxiv
Top 0.1%
18.7%
Show abstract

This evaluation of 36,000 clinical vignettes found that next-generation reasoning large language models, o3-mini and DeepSeek-R1, frequently perpetuate racial and gender stereotypes for common medical conditions, indicating that advancements in reasoning do not inherently improve representational fairness.

10
A Better Way: Initial Acceptability Testing of Using Artificial Intelligence Tools to Accelerate Development of Trauma Clinical Guidance

Zavala Wong, G.; Rosenauer, S.; Church, C.; Sherifali, D.; Racey, M.; Grider, K.; Moreno, A. N.; Lagrone, L. N.; The 2025 Design for Implementation (DFI) Authorship Group,

2025-08-24 surgery 10.1101/2025.08.20.25334097 medRxiv
Top 0.1%
18.7%
Show abstract

IntroductionRepresentatives of the trauma community have voiced a need for a new approach to developing clinical guidance. In this study, we test the initial acceptability of a proposed 12-step approach that aims to reduce the current clinical guidance timeline from more than 24 months to 24 weeks. MethodsInvestigators hypothesized that artificial intelligence (AI) tools could be leveraged to improve and make the process of clinical guidance development more efficient, facilitating AI initial output that could later be reviewed by subject matter experts (SMEs). Ensuring ethical standards and a collaborative design. Following the agile methodology, emphasizing continuous delivery and improvement, and the Practical, Robust Implementation and Sustainability Model (PRISM) framework, the investigators drafted a 12-step approach to clinical guidance development in 24 weeks. The process starts with the selection of a clinical topic and culminates in a bedside-ready clinical decision tree. ResultsThe 2025 Design for Implementation: The Future of Trauma Research & Clinical Guidance conference participants were invited to reflect on this new 12-step approach during two breakout sessions. Participants included a broad range of trauma providers, methodologists, patient representatives, technology, and marketing experts. Their recommendations highlighted: 1) multidisciplinary involvement, 2) need for resource-stratified recommendations, and 3) user-friendly features (offline and multilingual access). On a post conference survey (n=56), 64% were confident in AI accelerating the current development process. ConclusionsThe current landscape of clinical guidance offers significant opportunities for improvement. Key areas for enhancement include promoting collaboration across multiple disciplines and organizations, developing recommendations that consider resource variations, and utilizing new technologies, such as AI, to expedite the development process. This is crucial because ongoing delays lead to practices lagging behind current evidence. Further research is needed to rigorously test and refine how responsible use of AI can be integrated into expediting evidence integration into clinical guidance. Key MessagesO_ST_ABSWhat is already known on this topicC_ST_ABSCurrent clinical guidance typically takes 1-2 years to develop. Moreover, clinical guidance may not be published until a year or more after its completion, long after some recommendations become outdated, contributing to lagged evidence-informed practice. What this study addsThis study shares and tests the initial acceptability of a novel approach that aims to reduce the current clinical guidance timeline from 24 months to 24 weeks. It leverages existing artificial intelligence tools but with the critical input of subject matter experts (SMEs), ensuring ethical standards and collaborative design. SMEs shed light on critical steps and key areas that future clinical guidance needs to consider. How this study might affect research, practice or policyThe current landscape of clinical guidance offers significant opportunities for improvement. Key areas for enhancement include promoting collaboration across multiple disciplines and organizations, developing recommendations that consider resource variations, and utilizing new technologies, such as artificial intelligence, to expedite the development process.

11
Tension between timeliness and completeness of data in the initiation of cancer treatment: A qualitative study of oncology practice workflows and enduring health IT challenges

Samal, L.; Kyle, M. A.; Kilgallon, J. L.; Landrum, K. M.; Gawande, A. A.; Jacobson, J. O.; Hassett, M. J.

2025-05-21 health informatics 10.1101/2025.05.19.25324967 medRxiv
Top 0.1%
18.4%
Show abstract

IntroductionDiagnostic evaluation and treatment planning for newly diagnosed cancer requires a coordinated effort across multiple specialties. Delays in treatment initiation are common, leading to unnecessary anxiety and decreased survival. Given that timely treatment initiation is pivotal to providing high quality cancer care, we sought to characterize patient intake, workflows, and the role of health information technology (HIT) in a varied group of oncology practices nationwide. MethodsInterviews with oncologists were performed between March and September 2016, with follow-ups conducted between October and December 2021. Thematic analysis was used to assign codes to key elements of the transcripts, group these codes into conceptually distinct and clinically meaningful categories, and identify major cross-cutting themes. ResultsNine oncologists participated in an initial interview (one surgical, two radiation, six medical oncology). Four oncologists participated in a follow-up interview (one radiation, three medical oncology). In both time periods there was tremendous variation in staff roles and communication processes; some oncology practices obtained diagnostic studies before the first oncology consult visit, whereas others waited until after the initial consult visit to begin the diagnostic evaluation. Variability and tension were noted to arise from deficiencies in HIT, such as lack of interoperability, impaired speed and quality of data collection, cumbersome user interfaces, and variety of data types in oncology care. Oncologists reported only modest improvements in HIT between 2016 and 2021. ConclusionAssembling data to make a new cancer diagnosis and treatment plan is complex and time-intensive. HIT interoperability remains a quasi-manual process, contributing to preventable treatment delays. Federal policy supporting interoperability provides an opportunity to develop HIT that supports care coordination and patient-centered care, but effective implementation of such tools will be challenging within current workflows.

12
Auditing What Was Said: The Epistemic Promise and Limits of Ambient AI in Clinical Practice

Misrai, V.; Bruchon, A.; Campan, A.; Loubes, J. M.; Piau, A.

2026-04-30 health systems and quality improvement 10.64898/2026.04.29.26351997 medRxiv
Top 0.1%
18.4%
Show abstract

Traditional audit methods that rely on written records often miss the nuances of clinical reasoning that influence patient care. Ambient artificial intelligence captures spoken clinical encounters, allowing the analysis of real clinician-patient dialogue at scale. In a study of 124 urology consultations, a transcript-centered audit identified inter-physician variation and expert disagreement that conventional review missed. We explore the epistemic gains of this approach, its nonverbal blind spots, behavioral effects, technical vulnerabilities, and the EU AI Acts regulatory landscape.

13
Clinician Perceptions of Generative Artificial Intelligence Tools and Clinical Workflows: Potential Uses, Motivations for Adoption, and Sentiments on Impact

Ruan, E. L.; Alkattan, A.; Elhadad, N.; Rossetti, S. C.

2024-07-31 health informatics 10.1101/2024.07.29.24311177 medRxiv
Top 0.1%
18.2%
Show abstract

Successful integration of Generative Artificial Intelligence (AI) into healthcare requires understanding of health professionals perspectives, ideally through data-driven approaches. In this study, we use a semi-structured survey and mixed methods analyses to explore clinicians perceptions on the utility of generative AI for all types of clinical tasks, familiarity and competency with generative AI tools, and sentiments regarding the potential impact of generative AI on healthcare. Analysis of 116 clinician responses found differing perceptions regarding the usefulness of generative AI across clinical workflows, with information gathering from external sources rated highest and communication rated lowest. Clinician-generated prompt suggestions focused most often on clinician decision making and were of mixed quality, with participants more familiar with generative AI suggesting more high-quality prompts. Sentiments regarding the impact of generative AI varied, particularly regarding trustworthiness and impact on bias. Thematic analysis of open-ended comments highlighted concerns about patient care and the role of clinicians.

14
Patterns and Predictors of Artificial Intelligence Use Among Healthcare Professionals in the United States and United Kingdom: A Cross-National Survey

Sezgin, E.; Lee, J. A.; Jadczyk, T.; Taxter, A. J.

2026-05-06 health informatics 10.64898/2026.05.01.26352171 medRxiv
Top 0.1%
18.1%
Show abstract

ObjectiveWe surveyed 524 healthcare professionals (HCPs) in the United States and United Kingdom to examine workplace generative AI use, access, and barriers in two high-maturity health settings. MethodsThis cross-sectional survey compared AI usage breadth, access modes, and barriers among HCPs, stratified by country and professional role. ResultsOverall, 75.8% of HCPs reported recent AI use, mainly for documentation, literature search, and clinical decision support. Usage breadth was similar by country, but role differences were pronounced. Physicians reported broader use and were significantly more likely to access AI via personal, non-employer-provided tools (60.4% vs. 31.0% for nurses; P<.01). Personal tools were the most common access mode overall (40.1%). ConclusionAI use is common, but institutional access lags adoption. Shifting use from personal accounts toward governed, approved systems is a key priority.

15
Large Language Model Performance in UK Advice & Guidance: A Pilot Study in Neurology

Healy, J.; Marvasti, A.; Wallace, D.; Baheerathan, A.; Ghosh, A.; Kossoff, J.; Thio, S.; Balaratnam, M.; Haider, S.; Ellershaw, S.; Dobson, R.

2026-05-18 neurology 10.64898/2026.05.13.26353081 medRxiv
Top 0.1%
17.4%
Show abstract

Background: Large language models (LLMs) demonstrate strong performance in controlled medical environments such as multiple choice exams, but their utility in real-world clinical workflows remains unproven. The NHS Advice & Guidance (A&G) service, where Primary Care clinicians can submit text-based queries to specialists, provides an environment for evaluating the clinical performance of LLMs as a specialist. Methods: We compared responses from MedGemma 4B-IT, an open-weight model deployed locally on hospital infrastructure, against specialist neurologist responses across 50 adult neurology A&G cases from University College London Hospital. Two neurologists and two GPs rated 80 blinded and 20 unblinded responses for outcome, safety, efficacy, and feasibility using standardised criteria; outcome was a binary correct/incorrect, while other domains were scored 1-5. Inter-rater reliability was assessed using intraclass correlation coefficients. Results: Although there were no statistically significant differences between blinded specialist neurologists and LLM responses across any domain (outcome: 84% vs 82%, p=0.67; safety: 3.98 vs 4.02, p=0.85; efficacy: 4.06 vs 3.98, p=0.61; feasibility: 4.39 vs 4.20, p=0.45), 10% of LLM responses received concerning scores ([&le;]2 average score) compared to 0% of human responses, indicating potentially clinically important tail risk. Furthermore, unblinded results showed a preference for human responses, with human ratings being preferred across all domains. Only 51% of binary outcomes had unanimous agreement and inter-rater agreement was moderate across other domains (ICC 0.50-0.52). Conclusions: In this pilot study, aggregate scores between blinded human and LLM responses were similar, and no statistically significant differences were detected in this exploratory sample. However, aggregate metrics masked clinically important edge-case failures in LLM responses. Pronounced inter-rater variability and the potential impact of LLM/human syntax on blinded rater judgements highlight the challenges in establishing robust evaluation frameworks for clinical LLM deployment

16
A Retrospective Evaluation of the Microsoft Healthcare Agent Orchestrator for Tumor Board Patient Summaries

Roy, J.; Korleski, J. B.; Augustin, R. C.; Yefet, L.; Jensen, Z. D.; Ehman, E. C.; Zadeh, G.; Conners, A. L.; Tevaarwerk, A. J.; Korfiatis, P.

2026-06-01 health informatics 10.64898/2026.05.22.26353812 medRxiv
Top 0.1%
16.7%
Show abstract

Background: Preparing tumor board patient summaries is time intensive. Large-language-model based systems may automate summarization but require real-world evaluation prior to clinical use. We performed an exploratory retrospective evaluation of the Microsoft Healthcare Agent Orchestrator (HAO), deployed in a Mayo Clinic controlled staged environment, to generate tumor board-style patient summaries from retrospective Electronic Health Record (EHR) notes. Methods: HAO generated summaries for breast, hepatobiliary, and neuro-oncology tumor board cases using up to the most recent 1,000 clinical notes. Clinician reviewers evaluated outputs via REDCap surveys across perceived factuality, completeness, clarity/conciseness, temporal cohesion, comparative performance, safety, and clinical utility (0-4 Likert scale). Reviewers were permitted to query the HAO chat interface to address missing details. Automated factuality was assessed using TBFact (bidirectional entailment), reporting precision and recall against available reference summaries. Results: Among 57 survey responses from 5 different physicians, mean scores exceeded 2.8 across domains, with medians of 3 for most axes. In an exploratory comparison, oncology fellows required less time to review HAO-generated summaries than to manually generate patient summaries (mean difference 13.57 minutes per patient, p<0.001), although this difference may be influenced by prior familiarity with the same cases; 96% of survey responses indicated that HAO would save time. TBFact evaluations showed higher recall than precision across domains, consistent with broad capture of reference content alongside additional content that was not present in gold-standard summaries. Attribution was viewed favorably but showed issues with primary-source specificity and link reliability. Conclusions: In a controlled Mayo environment, HAO demonstrated moderate performance and was associated with reduced review time for tumor board preparation. These findings are promising but preliminary and do not establish clinical safety, noninferiority to manual review, or readiness for routine clinical use. Limitations, including verbosity, specialty-specific content gaps, and inconsistent attribution, highlight the need for iterative refinement and further evaluation.

17
An LLM-Based Comparison of Ambient AI Scribes for Clinical Documentation

Jain, J.; Kaan, J.; Jain, S.; Young, A.; Martinez, C.; Kartsonis, W.; Ortiz, C.; Cheng, R.; Jaklitsch, E.; Cherukuri, S.; Qilleri, A.; Tassiopoulos, A.

2025-06-26 primary care research 10.1101/2025.06.24.25330085 medRxiv
Top 0.1%
16.1%
Show abstract

Ambient AI scribes have become an increasingly promising option for automating clinical documentation, with dozens of enterprise solutions available. It remains uncertain whether models with domain-specific tuning outperform naive models "out of the box." This study evaluated five commercial AI scribes, alongside a custom solution using the base model of GPT-o1 without fine-tuning, as well as an experienced human scribe, in a series of simulated clinical encounters. Generated notes from these parties were scored by large language models (LLMs) using a rubric assessing completeness, organization, accuracy, complexity handling, conciseness, and adaptability. Our naive solution achieved scores comparable with industry-leading solutions across all rubric dimensions. These findings suggest that the added value of domain-specific training in ambient AI medical scribes may be limited when compared to base foundation models.

18
Optimising the Usability of AI Driven Augmented Reality Displays of Critical Structures During Surgery - An International Study of Surgeon-Computer Interaction

Ramirez Herrera, R.; Khan, D. Z.; Wijekoon, A.; Bano, S.; Clarkson, M. J.; Marcus, H.; Blandford, A.; CARES Evaluation Group,

2026-06-03 surgery 10.64898/2026.06.02.26354758 medRxiv
Top 0.1%
16.0%
Show abstract

In many endoscopic surgical procedures, the surgical team must identify and remove pathological tissue while avoiding critical structures such as arteries and nerves. Augmented reality (AR) offers potential support by overlaying visual information about the location of pathology and critical structures directly onto the operative field, enhancing spatial awareness and surgical navigation. However, limited research has evaluated how best to design and present AR overlays in ways that align with surgical workflow and perception. This study investigates surgeons' preferences across three key AR overlay dimensions: Design (how anatomy is visualised: outlines, heatmaps, masks, or centroids), Trigger (how and when overlays are activated: always visible, activated by the user, or triggered by instrument position), and Placement (where the overlay appears: above or below the surgical instrument). We take endoscopic pituitary adenoma surgery as a high-risk exemplar. Using a web-based prototype, 38 neurosurgeons ranked options and provided qualitative feedback. Surgeons preferred outline designs for clarity, user-activated triggers for control of information flow and distraction minimisation, and below-instrument placement for better spatial awareness. Preferences were consistent across experience levels and emphasised the importance of balancing visual saliency with cognitive load, to facilitate surgical navigation without distraction or disruption. These findings inform AR interface design, but require evaluation for impact on surgical performance and safety in further physical simulation and clinical studies.

19
Human-AI Collaboration in Clinical Reasoning: A UK Replication & Interaction Analysis

Healy, J.; Kossoff, J.; Lee, M.; Hasford, C.

2025-08-27 health informatics 10.1101/2025.08.25.25334383 medRxiv
Top 0.1%
15.5%
Show abstract

ObjectiveA paper from Goh et al found that a large language model (LLM) working alone outperformed American clinicians assisted by the same LLM in diagnostic reasoning tests [1]. We aimed to replicate this result in a UK setting and explore how interactions with the LLM might explain the observed gaps in performance. Methods and AnalysisThis was a within-subjects study of UK physicians. 22 participants answered structured questions on 4 clinical vignettes. For 2 cases physicians had access to an LLM via a custom-built web-application. Results were analysed using a mixed-effects model accounting for case difficulty and the variability of clinicians at baseline. Qualitative analysis involved coding of participant-LLM interaction logs and evaluating the rates of LLM use per question. ResultsPhysicians with LLM assistance scored significantly lower than the LLM alone (mean difference 21.3 percentage points, p < 0.001). Access to the LLM was associated with improved physician performance compared to using conventional resources (73.7% vs 66.3%, p = 0.001). There was significant heterogeneity in the degree of LLM-assisted improvement (SD 10.4%). Qualitative analysis revealed that only 30% of case questions were directly posed to the LLM, which suggests that under-utilisation of the LLM contributed to the observed performance gap. ConclusionWhile access to an LLM can improve diagnostic accuracy, realising the full potential of human-AI collaboration may require a focus on training clinicians to integrate these tools into their cognitive workflows and on designing systems that make these integrations the default rather than an optional extra.

20
REMOTE-Neuro: Co-produced Recommendations to Optimise Remote Neurology

Fuller, P.; Fearn, S.; Dace, S.; Wollam, A.; Zarkali, A.; Cowan, A.; Mountney, S.; Carr, G.; Eriksson, S. H.; Kipps, C.

2025-11-27 neurology 10.1101/2025.11.25.25340986 medRxiv
Top 0.1%
15.5%
Show abstract

ObjectiveTo examine stakeholder experiences of remote neurology outpatient care and to co-produce an evidence-based framework to support safe, equitable and sustainable service delivery. MethodsWe undertook an inductive thematic analysis of free-text responses from three national surveys: a patient and carer survey conducted by The Neurological Alliance (2021; n = 2,463) and two surveys of neurologists conducted by the Association of British Neurologists (2020 and 2021; n = 593). Findings were validated through co-production workshops and interviews with patients, carers and healthcare professionals in 2024 (n = 64). Themes were triangulated and refined to generate stakeholder-endorsed recommendations. ResultsParticipants valued flexible choice in consultation modality, recognising the accessibility and convenience of remote care, but expressed concerns about clinical quality, privacy and equity. Both patients and clinicians viewed remote care as a distinct skill set requiring tailored training and stronger digital infrastructure. Importantly, some participants perceived remote appointments as less legitimate than in-person consultations, a novel and under-recognised challenge with implications for engagement and health equity. Five domains aligned with NHS transformation principles were identified: (1) patient-centred care, (2) neurology and specialist area considerations, (3) clinical safety and quality, (4) capacity and sustainability, and (5) operational efficiency. These findings informed the REMOTE-Neuro Framework (REcommendations for optimising Modality, Operational efficiency, Training and Equity in NEUROlogy). ConclusionsRemote care continues to offer significant benefits to both patients and clinicians. However, five years post-pandemic, there are still unresolved issues which limit its effective integration with face-to-face (F2F) care. Practice implicationsREMOTE-Neuro provides the first co-produced set of recommendations to support safe, inclusive and sustainable remote neurology practice. Grounded in over 3,000 stakeholder perspectives and aligned with NHS transformation priorities, it offers a practical roadmap for implementation in neurology and a transferable model for other specialties. What is already known on this topicO_LIRemote consultations are now a routine component of neurology outpatient care. C_LIO_LIHowever, there is limited evidence on how to integrate remote and face-to-face (F2F) modalities safely, effectively, and equitably. C_LI What this study addsO_LIThis large, multi-dataset, co-produced study presents the first evidence-based national framework to optimise remote neurology services. C_LIO_LIIt identifies key patient, clinician and system-level factors that shape the effectiveness, safety and perceived value of remote neurology care. C_LI How this study might affect research, practice or policyO_LIThe REMOTE-Neuro framework provides actionable, co-produced recommendations that directly operationalise NHS outpatient transformation priorities. C_LIO_LIIt offers a practical structure to guide service design, clinical decision-making, training, digital inclusion strategies, and future evaluation of hybrid neurology models. C_LI